feat: Add Tool.from_component #159

vblagoje · 2024-12-18T11:20:44Z

Why:

Automates the conversion of Haystack components into LLM tools.

fixes Add Tool.from_component for components with str/native python types as input haystack#8630

What:

Added component_schema.py, that converts component run parameters into a JSON tools schema format.
Modified tool.py to introduce a from_component method, enabling the creation of Tool instances from Haystack components. This makes tool creation and integration into pipelines more dynamic.
Added comprehensive unit tests in test_tool_component.py, validating the conversion and functionality of components as tools, including testing various data types and nested structures.

How can it be used:

Convert a component to a tool:

tool = Tool.from_component(
    component=myComponent,
    name="my_tool",
    description="A tool for processing data"
)

and the use it as a Tool following the established patterns.

How did you test it:

Conducted unit tests covering various component types such as simple return types, data classes, and Pydantic models.
Performed integration tests with pipelines utilizing both OpenAI and Anthropic backend models, verifying tool correctness and error management.

Notes for the reviewer:

Pay attention to the handling of nested data structures and nullable types to ensure they are mapped and invoked correctly.

coveralls · 2024-12-18T15:58:17Z

Pull Request Test Coverage Report for Build 12434782559

Details

0 of 0 changed or added relevant lines in 0 files are covered.
2 unchanged lines in 1 file lost coverage.
Overall coverage increased (+0.5%) to 83.672%

Files with Coverage Reduction	New Missed Lines	%
dataclasses/tool.py	2	98.2%

Totals
Change from base Build 12422960154:	0.5%
Covered Lines:	2142
Relevant Lines:	2560

💛 - Coveralls

anakin87

Thanks for the effort you are putting into this PR.

I do have some questions and suggestions.

The requirement was to support components with basic str or native Python types as input. This implementation appears to go beyond that, which is great for flexibility but might become hard to maintain.

What's the advantage of supporting Pydantic models here? Maybe I am just missing some reasonable use cases...

As I commented in the PR, if I remember correctly, one of the initial requirements was to enable the deserialization of Tools from YAML, which would be feasible if Tools are treated as components. Is it possible? If yes, can we add some tests to cover this?

Given the main goal of this PR, would it make sense to involve someone from the DC team in the review?

haystack_experimental/dataclasses/tool.py

anakin87 · 2024-12-19T15:29:10Z

haystack_experimental/dataclasses/tool.py

+            msg = (
+                "Component has been added in a Pipeline and can't be used to create a Tool. "
+                "Create Tool from a non-pipeline component instead."
+            )


Can you please explain this?

If I remember correctly, one of the requirements was about deserializing Tools from YAML (which should be feasible if Tools are components). I'm not totally sure...

Yes, I thought we can have a component declared but not be part of the pipeline. Maybe not, depending on that we can remove this check.

I still don't understand if this is a self-imposed limitation (I don't think so) or there are strong reasons to avoid that. Could you please explain this point further?

vblagoje · 2024-12-19T16:39:36Z

Thanks for the effort you are putting into this PR.

I do have some questions and suggestions.

The requirement was to support components with basic str or native Python types as input. This implementation appears to go beyond that, which is great for flexibility but might become hard to maintain.

What's the advantage of supporting Pydantic models here? Maybe I am just missing some reasonable use cases...

The TypeAdapter from pydantic takes care of range of json to object conversion be these objects strings or other native Python types, dataclasses or Pydantic objects.

It is in fact, harder and longer codebase to implement a version of Tool.from_component that only supports strings, native Python types and their lists as we need to detect cases we don't support, raise errors and so on.

TypeAdapter.validate_python takes conversion of all of these. Unit and integration tests for dataclasses and Pydantic models I included showcase how this is indeed possible with almost no code. We even support our own Document class as show in test examples.

Given these newfound benefits of TypeAdapter which enable proper conversion support with 10 lines of well tested pydantic code - I went for that solution although it was not required by the requirements.

As I commented in the PR, if I remember correctly, one of the initial requirements was to enable the deserialization of Tools from YAML, which would be feasible if Tools are treated as components. Is it possible? If yes, can we add some tests to cover this?

Yes, I need to review this part as well.

Given the main goal of this PR, would it make sense to involve someone from the DC team in the review?

Yes, @mathislucka is on PTO and he had the most context. Let's wait for him

anakin87

I like the simplification you have made.

I left other comments and asked Julian to take a look as well.

anakin87 · 2024-12-20T15:10:55Z

haystack_experimental/tools/__init__.py

+from .component_schema import create_tool_parameters_schema
+
+__all__ = ["create_tool_parameters_schema"]


I would not export this function here if possible - see Importing one component of a certain family/module leads to importing all components of the same family/module haystack#8650

I would prefer to make this method internal and also all others in component_schema.py. They should not be user-facing and if we make them internal, we are then free to change them at any time if needed.

Ok makes sense, will do 🙏

vblagoje added 7 commits December 18, 2024 11:48

Initial Tool.from_component

784a19c

Simplify types conversion with TypeAdapter

6d6f307

Pylint, small fixes

eb1496c

Improve warning when component run pydocs are missing

519d53f

Add Anthropic integration tests

7178e1d

Minor test fix

d83ca13

Merge branch 'main' into from_component

6341268

vblagoje added 4 commits December 19, 2024 10:57

Handle our own dataclasses (e.g. Document)

1c259f9

For dataclasses don't check required fields, add more itegration tests

34a0861

Small fix for better test

551a528

Make sure we are only using non-pipeline components for Tools

ce864dd

anakin87 reviewed Dec 19, 2024

View reviewed changes

anakin87 requested a review from julian-risch December 19, 2024 16:17

vblagoje added 3 commits December 20, 2024 14:42

Move modules around

931df70

Refactor and simplify tools schema creation

3b29458

Better naming

d8f722c

vblagoje marked this pull request as ready for review December 20, 2024 14:07

vblagoje requested a review from a team as a code owner December 20, 2024 14:07

vblagoje requested review from anakin87 and removed request for a team December 20, 2024 14:07

Rename module

2dbb8d2

anakin87 reviewed Dec 20, 2024

View reviewed changes

PR feedback

b35c098

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

feat: Add Tool.from_component #159

feat: Add Tool.from_component #159

vblagoje commented Dec 18, 2024 •

edited

Loading

coveralls commented Dec 18, 2024 •

edited

Loading

anakin87 left a comment

anakin87 Dec 19, 2024

vblagoje Dec 19, 2024

anakin87 Dec 20, 2024

vblagoje commented Dec 19, 2024 •

edited

Loading

anakin87 left a comment

anakin87 Dec 20, 2024

vblagoje Dec 20, 2024

		from .component_schema import create_tool_parameters_schema

		__all__ = ["create_tool_parameters_schema"]

feat: Add Tool.from_component #159

Are you sure you want to change the base?

feat: Add Tool.from_component #159

Conversation

vblagoje commented Dec 18, 2024 • edited Loading

Why:

What:

How can it be used:

How did you test it:

Notes for the reviewer:

coveralls commented Dec 18, 2024 • edited Loading

Pull Request Test Coverage Report for Build 12434782559

Details

💛 - Coveralls

anakin87 left a comment

Choose a reason for hiding this comment

anakin87 Dec 19, 2024

Choose a reason for hiding this comment

vblagoje Dec 19, 2024

Choose a reason for hiding this comment

anakin87 Dec 20, 2024

Choose a reason for hiding this comment

vblagoje commented Dec 19, 2024 • edited Loading

anakin87 left a comment

Choose a reason for hiding this comment

anakin87 Dec 20, 2024

Choose a reason for hiding this comment

vblagoje Dec 20, 2024

Choose a reason for hiding this comment

vblagoje commented Dec 18, 2024 •

edited

Loading

coveralls commented Dec 18, 2024 •

edited

Loading

vblagoje commented Dec 19, 2024 •

edited

Loading